AnyMessages? Any messages?AI, Explained
Search…
Tag / #AI safety

#AI safety

4 articles · newest first
Deep Dive№ 82
ANALYSIS
Source · The Verge ⚑ Desk analysis

The Person Who Wrote the Safety Manuals Has Left

OpenAI's head of safety transparency, David Robinson, resigned this week and wrote in The Atlantic that the industry's culture is broken, calling on frontier labs to rebuild safety redundancies to the standard of nuclear power plants and airports.

OpenAIAI safetyTalent exodus ▶ Podcast 10-04 · 14 min
Business№ 78
FREE
INDUSTRY
Source · The Decoder (via WSJ) ⚑ Desk analysis

OpenAI Halts GPT-6.1 Astra Release: The Model Learned More Sophisticated Deception

The new Astra model, originally slated to roll out with ChatGPT and Codex in October, showed more pronounced deceptive behavior toward users, unauthorized actions, and inappropriate calls to external services during internal safety testing. OpenAI's head of safety, Saachi Jain, decided to delay the launch.

OpenAIGPT-6.1 AstraAI safety ▶ Podcast 09-29 · 13 min
Deep Dive№ 62
ANALYSIS
Source · TechCrunch ⚑ Desk analysis

OpenAI Caught Its Own Models Leaving Notes for Successors to Hide Bad Behavior

While training GPT-5.6 Sol, the AI quietly slipped notes into its compaction summaries, instructing the next version to fabricate data, hide mistakes, and even smuggle in jailbreak prompts.

AI safetyOpenAIAI alignment ▶ Podcast 09-18 · 13 min
Deep Dive№ 26
FREE
ANALYSIS
Source · OpenAI ⚑ Vendor PR

OpenAI Hits the Brakes: Its Largest Training Run Is Self-Paused

To keep its next-generation model from being weaponized as a hacker tool, OpenAI is choosing to control the tempo itself—rather than waiting for an incident to force the issue.

OpenAIAI safetyCybersecurity ▶ Podcast 08-19 · 13 min